NSF PAR Search | NSF Public Access Repository

Note: When clicking on a Digital Object Identifier (DOI) number, you will be taken to an external site maintained by the publisher. Some full text articles may not yet be available without a charge during the embargo (administrative interval).
What is a DOI Number?

Some links on this page may take you to non-federal websites. Their policies may differ from this site.

ActiveRIR: Active Audio-Visual Exploration for Acoustic Environment Modeling

Somayazulu, Arjun; Majumder, Sagnik; Chen, Changan; Grauman, Kristen (August 2024, International Conference on Intelligent Robots and Systems (IROS))

Full Text Available
Learning Spatial Features from Audio-Visual Correspondence in Egocentric Videos

Majumder, Sagnik; Al-Halah, Ziad; Grauman, Kristen (June 2024, IEEE Conference on Computer Vision and Pattern Recognition (CVPR))

Full Text Available
Few-Shot Audio-Visual Learning of Environment Acoustics

Majumder, Sagnik; Chen, Changan; Al-Halah, Ziad; Grauman, Kristen (January 2022, Advances in neural information processing systems)

Full Text Available
Move2Hear: Active Audio-Visual Source Separation

https://doi.org/10.1109/ICCV48922.2021.00034

Majumder, Sagnik; Al-Halah, Ziad; Grauman, Kristen (October 2021, IEEE/CVF International Conference on Computer Vision (ICCV))

We introduce the active audio-visual source separation problem, where an agent must move intelligently in order to better isolate the sounds coming from an object of interest in its environment. The agent hears multiple audio sources simultaneously (e.g., a person speaking down the hall in a noisy household) and it must use its eyes and ears to automatically separate out the sounds originating from a target object within a limited time budget. Towards this goal, we introduce a reinforcement learning approach that trains movement policies controlling the agent’s camera and microphone placement over time, guided by the improvement in predicted audio separation quality. We demonstrate our approach in scenarios motivated by both augmented reality (system is already co-located with the target object) and mobile robotics (agent begins arbitrarily far from the target object). Using state-of-the-art realistic audio-visual simulations in 3D environments, we demonstrate our model’s ability to find minimal movement sequences with maximal payoff for audio source separation.
more » « less
Full Text Available

Search for: All records